Papers with three-step process
ScholarBench: A Bilingual Benchmark for Abstraction, Comprehension, and Reasoning Evaluation in Academic Contexts (2025.findings-emnlp)
Copied to clipboard
| Challenge: | ScholarBench evaluates domain-specific knowledge of large language models (LLMs) prior benchmarks lack the scalability to handle complex academic tasks. |
| Approach: | ScholarBench evaluates the academic reasoning ability of large language models . the benchmark is constructed through a three-step process . |
| Outcome: | ScholarBench evaluates the academic reasoning ability of large language models . the benchmark comprises 5,031 examples in Korean and 5,309 examples in English . |
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark (2025.acl-long)
Copied to clipboard
Xiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang, Kai Zhang, Shengbang Tong, Yuxuan Sun, Botao Yu, Ge Zhang, Huan Sun, Yu Su, Wenhu Chen, Graham Neubig
| Challenge: | Recent advances in multimodal large language models have led to progress in tackling complex reasoning tasks that combine textual and visual information. |
| Approach: | They introduce a robust version of the Massive Multi-discipline Multimodal Understanding and Reasoning (MMMU) benchmark. |
| Outcome: | The proposed model performs lower on MMMU-Pro than on the previous benchmark, ranging from 16.8% to 26.9%. |
Opinions in Interactions : New Annotations of the SEMAINE Database (2022.lrec-1)
Copied to clipboard
| Challenge: | a new method for the detection of opinions in interactions is proposed . a dataset of dyadic interactions is annotated continuously in two affective dimensions related to the emotions . |
| Approach: | They propose to annotate opinions over a multimodal corpus of dyadic interactions . they use a d-acting algorithm to annnotate the opinions of a speaker . |
| Outcome: | The proposed method allows to obtain a precise annotation regarding the opinion of a speaker. |
PIRB: A Comprehensive Benchmark of Polish Dense and Hybrid Text Retrieval Methods (2024.lrec-main)
Copied to clipboard
| Challenge: | PIRB is a framework for text information retrieval in Polish . existing and new datasets are evaluated to evaluate the performance of 41 models . |
| Approach: | They propose a framework for 41 text information retrieval tasks in Polish . they evaluate over 20 dense and sparse retrieval models and build sparser-dense hybrid retrievers . |
| Outcome: | The proposed framework outperforms the best available methods in 41 tasks for Polish . the proposed models outperformed the best solutions available to date . |